Papers with human-robot interaction
Effects of Gender Stereotypes on Trust and Likability in Spoken Human-Robot Interaction (L18-1)
Copied to clipboard
| Challenge: | a study investigates the influence of gender stereotypes on trust and likability of humanoid robots . explicit gender and stereotypicality of a task are manipulated to influence robot behavior . future research may look into situational variables that drive stereotypification in robot interaction . |
| Approach: | They investigated the influence of gender stereotypes on trust and likability of robots . they used explicit (name and voice) and implicit (personality) genders to manipulate stereotypical tasks . future research may look into situational variables that drive stereotypization . |
| Outcome: | The findings suggest that gender stereotypes need to be differentiated in robot interaction . the gender and personality characteristics of robots influence trust and likability . |
Dialogue-AMR: Abstract Meaning Representation for Dialogue (2020.lrec-1)
Copied to clipboard
Claire Bonial, Lucia Donatelli, Mitchell Abrams, Stephanie M. Lukin, Stephen Tratz, Matthew Marge, Ron Artstein, David Traum, Clare Voss
| Challenge: | Abstract Meaning Representation (AMR) does not capture the illocutionary force or speaker’s intended contribution in the broader dialogue context. |
| Approach: | They propose a schema that enriches Abstract Meaning Representation (AMR) it provides a semantic representation for facilitating Natural Language Understanding (NLU) in dialogue systems. |
| Outcome: | The proposed schema provides a semantic representation for facilitating Natural Language Understanding (NLU) in human-robot dialogue systems. |
Learning Physical Common Sense as Knowledge Graph Completion via BERT Data Augmentation and Constrained Tucker Factorization (2020.emnlp-main)
Copied to clipboard
| Challenge: | Physical commonsense learning is an essential part of human-robot interaction . existing methods of learning physical commons sense suffer from generalization . |
| Approach: | They propose to use physical commonsense learning as a knowledge graph completion problem to better use latent relationships among training samples. |
| Outcome: | The proposed method outperforms existing methods in the human-robot interaction problem. |
Aligning Images and Text with Semantic Role Labels for Fine-Grained Cross-Modal Understanding (2022.lrec-1)
Copied to clipboard
| Challenge: | Currently, image retrieval systems can retrieve relevant results for diverse inputs, but they do not provide a way to intentionally inject variety into the search results. |
| Approach: | They propose a multimodal dataset that combines semantic annotations with image bounding boxes. |
| Outcome: | The proposed system improves image retrieval performance and flexibility. |
FARMI: A FrAmework for Recording Multi-Modal Interactions (L18-1)
Copied to clipboard
Patrik Jonell, Mattias Bystedt, Per Fallgren, Dimosthenis Kontogiorgos, José Lopes, Zofia Malisz, Samuel Mascarenhas, Catharine Oertel, Eran Raveh, Todd Shore
| Challenge: | a new framework for recording multi-modal data is needed to capture multi-party, richly recorded corpora and perform real-time processing of such data. |
| Approach: | They propose an open-source processing architecture for corpora and real-time processing . they deploy the architecture in a multi-party deception game with six humans and one robot . |
| Outcome: | The proposed architecture is agnostic to hardware and programming languages, although it's mostly written in Python. |
GazeVQA: A Video Question Answering Dataset for Multiview Eye-Gaze Task-Oriented Collaborations (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies on the use of exocentric and egocentric videos in video question answering are focusing on eye-gaze information. |
| Approach: | They propose a task-oriented VQA dataset that captures eye-gaze information . they propose assisting models that ground the perceptual input into semantic information based on three different answer types . |
| Outcome: | The proposed model can ground the perceptual input into semantic information while reducing ambiguities. |
How Much Do Robots Understand Rudeness? Challenges in Human-Robot Interaction (2024.lrec-main)
Copied to clipboard
| Challenge: | This paper examines the pressing need to understand and manage inappropriate language within the evolving human-robot interaction landscape. |
| Approach: | They propose to use data cleaning methods to identify inappropriate language in real-time interactions and evaluate natural language models for their proficiency in discerning rudeness. |
| Outcome: | The proposed methods identify and mitigate inappropriate language in real-time interactions and evaluate natural language models for their proficiency in discerning rudeness. |
CityNavAgent: Aerial Vision-and-Language Navigation with Hierarchical Semantic Planning and Global Memory (2025.acl-long)
Copied to clipboard
Weichen Zhang, Chen Gao, Shiquan Yu, Ruiying Peng, Baining Zhao, Qian Zhang, Jinqiang Cui, Xinlei Chen, Yong Li
| Challenge: | Existing ground VLN agents struggle in aerial VLLN due to the lack of predefined navigation graphs and the exponentially expanding action space in long-horizon exploration. |
| Approach: | They propose a large language model-empowered aerial VLN agent that decomposes the long-horizon task into sub-goals with different semantic levels. |
| Outcome: | The proposed method achieves state-of-the-art performance with significant improvement in continuous city environments. |